02_COMPUTER_SCIENCE_NOTES (AI SUBSYSTEM)
Week 7 AI Subsystem · Dual Sovereign Core (AR / EN)
⚑ TOKENIZATION, EMBEDDING VECTORS & ATTENTION MECHANISMS
AYMAN ELMASRY
Computational Creative Director · AI Prompt Engineer
Founder of Ayman Elmasry LLC
πŸ”’ ⚑ AEL Sovereign Seal (Active Master Verification)
{
  "ael_seal": "AEL CS Encyclopedia β€” Β© Ayman Elmasry",
  "owner": "Ayman Elmasry",
  "legal_entities": [
    "Ayman Elmasry LLC (UAE)",
    "Ayman Elmasry Advertising & Marketing (Egypt)"
  ],
  "syllabus_source": "Harvard CS50x 2026-2027",
  "domain": "Week 7 (AI Subsystem): Tokenization, Embedding Vectors & Attention Mechanisms",
  "document_type": "02_Computer_Science_Notes",
  "methodology": "8-Stage Sub-Silicon Execution Paradigm",
  "system_version": "v3.0"
}

Computer Science Notes: Mathematical Underpinnings of LLMs (Tokenization & Vectors)

Textual Discretization (Tokenization)

To comprehend Large Language Models (LLMs) beyond superficial perceptions of artificial sentience, one must examine the rigorous mathematical transformations executing within high-dimensional vector spaces.

  • Token Decomposition: Words and subwords are mapped to discrete numerical token IDs. A single word may constitute one token or split into smaller subword chunks.
  • Context Window: The rigid, upper-bounded memory buffer defining the maximum sequence of past tokens the attention mechanism can retain to compute the subsequent probability matrix.

Vector Embeddings & Semantic Topography

  • High-Dimensional Vector Representation: Every token ID is transformed into a dense floating-point vector tensor (frequently spanning 1536, 4096, or more dimensions).
  • Semantic Proximity (Euclidean & Cosine Similarity): Conceptually synonymous or related entities cluster geometrically adjacent to one another within this multi-dimensional space.
===================================================================================
                   HIGH-DIMENSIONAL EMBEDDING VECTOR SPACE
===================================================================================

       [ Cat ] (0.25, 0.89, 0.12)  <── Close Cosine Distance ──> [ Dog ] (0.28, 0.85, 0.15)
       
                                   <── Massive Distance ────> [ Car ] (0.91, 0.05, 0.88)

===================================================================================

The Self-Attention Mechanism

Allows the model to dynamically compute query, key, and value (Q, K, V) correlation matrices between every token and all surrounding tokens, enabling flawless resolution of complex syntactic dependencies and ambiguous pronouns.